Papers by Sean Timothy Okonsky
PEaCE: A Chemistry-Oriented Dataset for Optical Character Recognition on Scientific Documents (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing open-source OCR models focus on scientific texts or generic printed English . Nougat is unable to parse tables in PubMed articles . |
| Approach: | They propose to train OCR models for scientific or generic printed English . Nougat is a popular tool for parsing academic documents, but unable to parse PubMed tables . |
| Outcome: | The proposed models perform better when trained on real-world records than those trained on synthetic records. |